Skip to main content

CLI Tools

Health data engineering happens in the terminal more often than the slides suggest. These are the command-line tools that repay learning.


FHIR​

HL7 FHIR Validator — the reference validator, run against resources, profiles and implementation guides. The fastest way to settle an argument about whether a resource is conformant.

java -jar validator_cli.jar patient.json -version 4.0.1 -ig hl7.fhir.us.core

SUSHI — compiles FHIR Shorthand (FSH) into profiles, extensions and value sets. Writing profiles as text you can diff and review beats editing JSON by hand.

sushi build .

IG Publisher — builds a full implementation guide site from FSH output and narrative pages.

fhir-py / fhirclient / fhir.resources — Python libraries with CLI-friendly entry points for scripted resource work.


HTTP and APIs​

curl — the baseline. If it does not work in curl, it does not work.

curl -s -H "Accept: application/fhir+json" \
"https://example.org/fhir/metadata" | jq '.rest[0].resource[].type'

httpie — friendlier syntax for exploratory work.

jq — the essential JSON processor: filter, reshape, extract.

jq '.entry[].resource | {id, birthDate}' bundle.json

yq — the same idea for YAML.

See APIs for monitoring and observability tooling.


Data wrangling​

csvkit — csvlook, csvstat, csvsql for inspecting and querying CSV extracts without opening a spreadsheet.

Miller (mlr) — CSV/TSV/JSON processing with named fields; excellent for reshaping export files.

DuckDB — SQL over CSV and Parquet files directly from the shell, fast enough for most analysis of routine health data extracts.

duckdb -c "SELECT district, count(*) FROM 'facilities.csv' GROUP BY 1"

pandas via a short script — when the transformation stops fitting on one line.


Databases and deployment​

  • psql — PostgreSQL client; \copy, \timing and EXPLAIN ANALYZE are the workhorses of query tuning.
  • pg_dump / pg_restore — backups and environment refreshes.
  • docker and docker compose — see Docker.
  • rclone, rsync — moving extracts and backups between environments.

Working safely​

  • Never pass credentials as command-line arguments. They land in shell history and process listings; use environment variables or credential files with restricted permissions.
  • De-identify before you experiment. Pull synthetic or de-identified data for exploratory work — see health data.
  • Script it once you have run it twice. Repeated manual pipelines drift.
  • Keep scripts in version control, including the ad-hoc ones. Today's quick fix is next quarter's undocumented dependency.